Papers with reward design

6 papers
Reasoning Structure Matters for Safety Alignment of Reasoning Models (2026.acl-long)

Copied to clipboard

Challenge: Large reasoning models (LRMs) achieve strong performance on complex reasoning tasks but often generate harmful responses to malicious user queries.
Approach: They propose a method that alters the reasoning structure of large reasoning models to achieve effective safety alignment by supervised fine-tuning.
Outcome: The proposed method is practical and generalizable, requiring no complex training or reward design.
A Fairness-Driven Method for Learning Human-Compatible Negotiation Strategies (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in AI and NLP have led researchers to develop techniques to build autonomous agents which can achieve human-level performance in bargaining games such as Deal-orno-Deal.
Approach: They propose a negotiation framework which incorporates fairness into reward design and search to learn human-compatible negotiation strategies.
Outcome: The proposed framework achieves more egalitarian negotiation outcomes and improves negotiation quality.
NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Neologism-aware machine translation aims to translate source sentences containing neologismes into target languages.
Approach: They propose an agentic framework for neologism-aware machine translation equipped with a Wiktionary-based search toolkit.
Outcome: The proposed framework is based on a Wiktionary-based search toolkit and a dedicated dataset for neologism-aware machine translation.
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in text-only "slow thinking" reasoning have prompted efforts to transfer this capability to vision-language models (VLMs).
Approach: They propose a VRM Reflection-V which enhances visual reflection based on reasoning data for cold-start and reward design for reinforcement learning.
Outcome: The proposed model improves visual reflection for cold-start and reward design for reinforcement learning (RL) it maintains a stronger and more consistent reliance on visual information during visual reasoning, indicating effective enhancement in visual reflection capabilities.
MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: Large-scale reinforcement learning (RL) methods have proven effective in enhancing the reasoning abilities of large language models.
Approach: They propose an open-source adaptation of the R1-Zero RL framework for machine translation (MT) their code is available at https://github.com/fzp0424/MT-R1-zero.
Outcome: The proposed framework surpasses towerinstruct-7B-v0.2 on the english-chinese benchmark by 1.26 points.
NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in Graphical User Interface (GUI) and embodied navigation have driven progress, yet these domains have largely evolved in isolation, with disparate datasets and training paradigms.
Approach: They propose a visual-target trajectory collection pipeline that generates trajectories for GUI and embodied tasks using a single formulation.
Outcome: The proposed agent outperforms state-of-the-art agents in GUI navigation, spatial affordance prediction, and embodied navigation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations